Papers with shortcut learning
DISCOSQA: A Knowledge Base Question Answering System for Space Debris based on Program Induction (2023.acl-industry)
Copied to clipboard
| Challenge: | a system that can answer complex natural language queries is developed for the European Space Agency . space debris are uncontrolled artificial objects left in orbit during normal operations or due to malfunctions . |
| Approach: | They propose a query-based system that can answer queries in natural language . it generates a program sketch from a natural language question and executes it against the database . |
| Outcome: | The proposed system can answer queries in natural language based on a natural language question generated by a query program . the system reduces overfitting and shortcut learning even with limited training data, the authors say . |
Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing validation studies assess output-level agreement; none test causal structure of LLM decisions mirrors that of human decisions. |
| Approach: | They compare topic sentiment and ideology labels using human annotators . they say this is evidence of shortcut learning by fine-tuning on ideology-labeled data . |
| Outcome: | The proposed model internalises a spurious sentiment–ideology coupling not operative in human judgment for this task. |
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs (2026.acl-short)
Copied to clipboard
| Challenge: | Existing multimodal reasoning models lack generalized spatial intelligence, a new study shows . a critical gap exists in the field of vision-centric reasoning, the authors argue . |
| Approach: | They evaluate 16 multimodal reasoning models using Chain-of-Though (CoT) based thinking . they find that CoT prompting consistently degrades performance in visual spatial reasoning . |
| Outcome: | The proposed model hallucinates visual details from textual priors even when the image is absent. |
Threat Scenarios and Best Practices to Detect Neural Fake News (2022.coling-1)
Copied to clipboard
| Challenge: | During the COVID-19 pandemic, inaccurate information made it hard for people to find reliable guidance when they needed it. |
| Approach: | They propose to use pretrained language models to generate fluent, original text . they argue that strong detectors should be released along with new generators . |
| Outcome: | The proposed system is prone to shortcut learning and should be released along with new generators. |
Language Prior Is Not the Only Shortcut: A Benchmark for Shortcut Learning in VQA (2022.findings-emnlp)
Copied to clipboard
Qingyi Si, Fandong Meng, Mingyu Zheng, Zheng Lin, Yuanxin Liu, Peng Fu, Yanan Cao, Weiping Wang, Jie Zhou
| Challenge: | Visual Question Answering (VQA) models are prone to learn the shortcut solution formed by dataset biases rather than the intended solution. |
| Approach: | They propose a dataset that considers varying types of shortcuts by constructing different distribution shifts in multiple OOD test sets. |
| Outcome: | The proposed dataset considers varying types of shortcuts by constructing different distribution shifts in multiple OOD test sets. |
An Investigation of LLMs’ Inefficacy in Understanding Converse Relations (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks for Large Language Models (LLMs) follow the data distribution of pre-training data. |
| Approach: | They propose a benchmark ConvRe focusing on converse relations which contains 17 relations and 1240 triples extracted from popular knowledge graph completion datasets. |
| Outcome: | The proposed benchmark focuses on converse relations, which contains 17 relations and 1240 triples extracted from popular knowledge graph completion datasets. |
Implicit Reasoning in Transformers is Reasoning through Shortcuts (2025.findings-acl)
Copied to clipboard
| Challenge: | Language models can perform step-by-step reasoning and achieve high accuracy in both in-domain and out-of-domain tests via implicit reasoning. |
| Approach: | They train GPT-2 from scratch on a curated multi-step mathematical reasoning dataset and conduct analytical experiments to investigate how language models perform implicit reasoning in multi- step tasks. |
| Outcome: | The proposed model performs better on multi-step tasks than the explicit reasoning model. |
Supervised and Unsupervised Probing of Shortcut Learning: Case Study on the Emergence and Evolution of Syntactic Heuristics in BERT (2025.findings-acl)
Copied to clipboard
| Challenge: | Contemporary language models (LMs) rely on shortcut learning, using superficial cues that are spuriously correlated with labels. |
| Approach: | They propose to use syntactic heuristics to learn shortcuts in BERT when performing a task in Natural Language Understanding to investigate where these shortcuts emerge, how they evolve and how they impact the latent knowledge of the LM. |
| Outcome: | The proposed model rely on syntactic heuristics when performing a task in Natural Language Understanding. |
Mitigating Shortcut Learning via Smart Data Augmentation based on Large Language Model (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods to improve shortcut learning performance are limited by manual definition of shortcuts and inherent confirmation bias during model training. |
| Approach: | They propose a method of Smart Data Augmentation based on Large Language Models to identify shortcuts and generate their anti-shortcut counterparts. |
| Outcome: | The proposed method shows an improvement of 5.61% across various natural language processing tasks. |
Data Drives Unstable Hierarchical Generalization in LMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Early in training, LMs can behave like n-gram models but eventually learn tree-based syntactic rules and generalize out of distribution (OOD). |
| Approach: | They study how complex data drives hierarchical rules, while less complex encourages shortcut learning . they find a model uses rules to generalize if its training data is *diverse* . |
| Outcome: | The proposed model learns to generalize hierarchically if its training data is complex . a model learn if it includes center-embedded clauses, a special syntactic structure . |
Exploring and Mitigating Shortcut Learning for Generative Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent large language models (LLMs) have incredible instruction-following capabilities while maintaining strong task completion ability. |
| Approach: | They propose a framework to encourage LLMs to Forget Spurious correlations and Learn from In-context information. |
| Outcome: | The proposed framework can mitigate shortcut learning by forging spurious correlations and learning from in-context information. |
Learning by Analogy: Diverse Questions Generation in Math Word Problem (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for solving math word problem (MWP) use shortcut learning to train solvers based on samples with a single question. |
| Approach: | They propose to generate diverse yet consistent questions from a common scenario . they then feed the equations to a question generator to obtain the diverse questions . their method leads to performance improvement on the current benchmark Math23K . |
| Outcome: | The proposed method generates diverse yet consistent questions with a variety of equations and questions . it improves on the current benchmark, which is based on the proposed method . |
HuaSLIM: Human Attention Motivated Shortcut Learning Identification and Mitigation for Large Language models (2023.findings-acl)
Copied to clipboard
| Challenge: | Large language models tend to rely on shortcut features that spuriously correlate with labels for prediction, which weakens their generalization on out-of-distribution samples. |
| Approach: | They propose a human attention guided approach to identifying shortcut learning that encourages the LLM-based target model to learn relevant features by exploring both human and neural attention. |
| Outcome: | The proposed approach improves the robustness of large language models on out-of-distribution (OOD) samples while not affecting the performance on IID data. |
Efficient Overshadowed Entity Disambiguation by Mitigating Shortcut Learning (2024.emnlp-main)
Copied to clipboard
Panuthep Tasawong, Peerat Limkonchotiwat, Potsawee Manakul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, Sarana Nutanong
| Challenge: | Entity disambiguation (ED) is crucial in natural language processing tasks such as question-answering and information extraction. |
| Approach: | They propose a method to reduce computational overhead on overshadowed entities by addressing shortcut learning. |
| Outcome: | The proposed method achieves state-of-the-art performance without compromising inference speed. |
The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) exhibit a critical failure mode: they exhibit brittle reasoning capabilities on out-of-distribution tasks. |
| Approach: | They propose a framework bridging Structural Causal Models and the Information Bottleneck principle to explain this paradox. |
| Outcome: | The proposed framework bridges the framework between SCM and IB principles to explain the problem. |
Semformer: Transformer Language Models with Semantic Planning (2024.emnlp-main)
Copied to clipboard
| Challenge: | Neural language models (LLMs) employ teacher forcing to predict tokens based on preceding ground truth tokens. |
| Approach: | They propose a method for training a Transformer language model that explicitly models the semantic planning of response. |
| Outcome: | The proposed method exhibits near-perfect performance and mitigates shortcut learning. |
NiuTrans.LMT: Toward Inclusive and Scalable Multilingual Machine Translation with LLMs (2026.acl-long)
Copied to clipboard
Yingfeng Luo, Ziqiang Xu, Yuxuan Ouyang, MuRun Yang, DingYang Lin, Kaiyan Chang, Tong Zheng, Bei Li, Peinan Feng, Quan Du, Tong Xiao, JingBo Zhu
| Challenge: | Large language models have significantly advanced Multilingual Machine Translation (MMT) yet scaling to many languages while maintaining robust performance across directions remains challenging. |
| Approach: | They propose a strategy to reduce the number of translations in one direction . they propose auxiliary parallel sentences to promote cross-lingual transfer . |
| Outcome: | The proposed model performs on par with or better than substantially larger baselines. |
Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement Learning (2026.acl-long)
Copied to clipboard
Zhaoyan Gong, Zhiqiang Liu, Songze Li, Xiaoke Guo, Yuanxiang Liu, Xinle Deng, Zhizhen Liu, Lei Liang, Huajun Chen, Wen Zhang
| Challenge: | Existing methods rely on fixed workflows and expensive closed-source APIs, limiting flexibility and scalability. |
| Approach: | They propose a temporal reasoning agent that trains on difficult questions first . they expand the action space with specialized internal actions alongside external action . |
| Outcome: | The proposed agent improves 19.8% over baselines on complex questions and multi-tasks. |
LLMs in Sarcasm Detection? It’s elementary! (Or is it?) (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are often cited for their sophisticated pragmatic reasoning, but they collapse to random guessing on organic human speech. |
| Approach: | They propose that LLMs have near-human competence in sarcasm detection . authors propose that this proficiency may be deceptive . |
| Outcome: | The proposed model performance on synthetic leaderboards is a statistical mirage of competence. |
CO-EVO: Co-evolving Semantic Anchoring and Style Diversification for Federated DG-ReID (2026.acl-long)
Copied to clipboard
| Challenge: | Existing frameworks for person re-identification fail to provide global supervision . stylistic gaps in the model can lead to shortcut learning . |
| Approach: | They propose a framework that aims to generalize a person's identity across multiple decentralized domains. |
| Outcome: | The proposed framework achieves state-of-the-art (SOTA) performance . it can generalize to unseen target environments without compromising privacy . |
A2O: LLM-based Agentic Learning of Action-to-Object Features for Video Action Recognition (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent action recognition based on vision–language pretraining and self-supervised video foundation models tends to induce spurious correlations and shortcut learning by relying on action-irrelevant cues. |
| Approach: | They propose a framework in which an LLM agent integrates the two approaches within an agentic learning paradigm to design motion features tailored to the target actions. |
| Outcome: | The proposed model is based on the commonsense knowledge of large language models (LLMs) and the open vocabulary object detector to make the model attend to objects in a video required for recognizing the target actions. |
From What Is Said to Why It Is Framed: Intent-Aware News Video Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing verification methods for short-form news videos neglect communicative intent . stylistic presentation and factual manipulation are often intertwined, resulting in shortcut learning . |
| Approach: | They propose a theory-grounded representation of communicative intent that captures creator stance, audience need activation, and communication strategy. |
| Outcome: | The proposed framework captures creator stance, audience need activation, and communication strategy. |